Goto

Collaborating Authors

 Wisconsin


Anyone can fake a scientific image with AI, tricking even academic journals – and undermining trust in science

AIHub

A photograph of Earth glowing in deep space, the Moon's cratered horizon stretching across its foreground, caught many people's eyes in April 2026. Astronauts captured the image while aboard NASA's Artemis II mission, and like the famous Apollo 8 "Earthrise" image, the picture felt instantly real and inspiring for many. But when almost anyone can fabricate a visually similar image in seconds from a text prompt using artificial intelligence, how do people decide which image is real? The proliferation of AI-generated science images in public spaces is not simply a misinformation problem. As a researcher who studies visual science communication and public trust, I believe it also contributes to a crisis of trust in science in the age of AI, and the tools scientists have long relied on to establish visual credibility are losing their grip.


Legendary WW2 submarine heads to Wisconsin for major facelift

Popular Science

The USS Silversides sank at least 23 ships in the Pacific theater. More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. The USS Silversides is one of the most decorated ships in US history. Breakthroughs, discoveries, and DIY tips sent six days a week. By signing up, you confirm you are 16+, will receive newsletters and promotional content and agree to our Terms of Use and acknowledge the data practices in our Privacy Policy .


The A.I. Gender Gap Meets the Parenting Gender Gap

The New Yorker

Women use A.I. less than men and do more of the cognitive work at home. The A.I. "family assistant" promises to bridge both divides. In a promotional video for Ollie, an A.I. family assistant, one of the company's founders, Bill Lennon, is about to announce his new venture in front of a camera crew when his phone starts pinging. "THE BABY JUST ON THE CARPET / like A LOT / WHILE I WAS ON ZOOM W AN INVESTOR." Then, a somewhat anticlimactic follow-up: "My sister coming for dinner btw did I tell you???" Lennon does not respond directly to these messages, which are presumably from his wife.


Offline Actor-Critic for Average Reward MDPs

Neural Information Processing Systems

We study offline policy optimization for infinite-horizon average-reward Markov decision processes (MDPs) with large or infinite state spaces. Specifically, we propose a pessimistic version of actor-critic methods using a computationally efficient linear function class for value function estimation. At the core of our method is a critic that computes a pessimistic estimate of the average reward under the current policy, as well as the corresponding policy gradient, by solving a fixedpoint Bellman equation, rather than solving a successive sequence of regression problems as in finite horizon settings. Due to the nature of our policy-based method, the critic only needs to solve a linear optimization problem with convex quadratic constraints. We show that a very mild data coverage requirement is sufficient for our algorithm to achieve O(ε 2) sample complexity for learning a near-optimal policy up to model misspecification errors. To our knowledge, this is the first result with optimal εdependence in the offline average reward setting.


Optimal Mistake Bounds for Transductive Online Learning

Neural Information Processing Systems

We resolve a 30-year-old open problem concerning the power of unlabeled data in online learning by tightly quantifying the gap between transductive and standard online learning. In the standard setting, the optimal mistake bound is characterized by the Littlestone dimension dof the concept class H(Littlestone, 1987). We prove that in the transductive setting, the mistake bound is at least Ω d . This constitutes an exponential improvement over previous lower bounds of Ω(loglog(d)), Ω p log(d), and Ω(log(d)), due respectively to Ben-David, Kushilevitz, and Mansour (1995, 1997), and Hanneke, Moran, and Shafer (2023). We also show that this lower bound is tight: for every d, there exists a class of Littlestone dimension d with transductive mistake bound O d . Our upper bound also improves upon the best known upper bound of (2/3) d from Ben-David et al. (1997). These results establish a quadratic gap between transductive and standard online learning, thereby highlighting the benefit of advance access to the unlabeled instance sequence. This contrasts with the PAC setting, where transductive and standard learning exhibit similar sample complexities.



OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Neural Information Processing Systems

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization, handwritten content extraction, and logical reasoning, remain underexplored. To bridge this gap, we introduce OCRBench v2, a large-scale bilingual text-centric benchmark with currently the most comprehensive set of tasks (4 more tasks than the previous multi-scene benchmark OCRBench), the widest coverage of scenarios (31diverse scenarios), and thorough evaluation metrics, with 10,000human-verified questionanswering pairs and a high proportion of difficult samples. Moreover, we construct a private test set with 1,500 manually annotated images. The consistent evaluation trends observed across both public and private test sets validate the OCRBench v2's reliability. After carefully benchmarking state-of-the-art LMMs, we find that most LMMs score below 50 (100 in total) and suffer from five-type limitations, including less frequently encountered text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning.


Global Minimizers of ℓp-Regularized Objectives Yield the Sparsest ReLU Neural Networks

Neural Information Processing Systems

Overparameterized neural networks can interpolate a given dataset in many different ways, prompting the fundamental question: which among these solutions should we prefer, and what explicit regularization strategies will provably yield these solutions? This paper addresses the challenge of finding the sparsest interpolating ReLU network--i.e., the network with the fewest nonzero parameters or neurons--a goal with wide-ranging implications for efficiency, generalization, interpretability, theory, and model compression. Unlike post hoc pruning approaches, we propose a continuous, almost-everywhere differentiable training objective whose global minima are guaranteed to correspond to the sparsest singlehidden-layer ReLU networks that fit the data. This result marks a conceptual advance: it recasts the combinatorial problem of sparse interpolation as a smooth optimization task, potentially enabling the use of gradient-based training methods. Our objective is based on minimizing ℓp quasinorms of the weights for 0 < p < 1, a classical sparsity-promoting strategy in finite-dimensional settings. However, applying these ideas to neural networks presents new challenges: the function class is infinite-dimensional, and the weights are learned using a highly nonconvex objective. We prove that, under our formulation, global minimizers correspond exactly to sparsest solutions. Our work lays a foundation for understanding when and how continuous sparsity-inducing objectives can be leveraged to recover sparse networks through training.


Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning

Neural Information Processing Systems

Large Language Models (LLMs) are increasingly used to simulate human users in interactive settings such as therapy, education, and social role-play. While these simulations enable scalable training and evaluation of AI agents, off-the-shelf LLMs often drift from their assigned personas, contradict earlier statements, or abandon role-appropriate behavior. We introduce a unified framework for evaluating and improving persona consistency in LLM-generated dialogue. We define three automatic metrics--prompt-to-line consistency, line-to-line consistency, and Q&A consistency--that capture different types of persona drift and validate each against human annotations. Using these metrics as reward signals, we apply multiturn reinforcement learning to fine-tune LLMs for three user roles: a patient, a student, and a social chat partner. Our method reduces inconsistency by over 55%, resulting in more coherent, faithful, and trustworthy simulated users.


Pioneering UK Nerve Lab harnesses AI to map effect of children's screen time

The Guardian

Tim Smith: 'Today's short-form, fast-paced, highly captivating content may affect children's attention, comprehension and emotional response'. Tim Smith: 'Today's short-form, fast-paced, highly captivating content may affect children's attention, comprehension and emotional response'. Pioneering UK Nerve Lab harnesses AI to map effect of children's screen time P arents are constantly being told to limit their children's screen time. A relatively slow-paced programme such as Bluey offers a very different viewing experience to a fast-moving action series such as PAW Patrol, yet both are broadly considered suitable for young children. This challenge is growing as the type of content children are exposed to evolves.